This is a placeholder. Final title will be filled later
نویسندگان
چکیده
This paper describes some experiments with pronunciation modeling for spontaneous speech in European Portuguese. The transducer framework provides an elegant way to combine a pronunciation lexicon of canonical forms with alternative pronunciation rules. The main phonological aspects that the rules are intended to cover are: vowel devoicing, deletion and coalescence, voicing assimilation, and simplification of consonantal clusters, both within words and across word boundaries. Our aligner proved sufficiently robust to be able to process fairly long dialogs with overlapping turns, despite many limitations, namely in terms of absence of models for voice quality changes.
منابع مشابه
This is a placeholder. Final title will be filled later
This paper presents a model to predict the phrase commands of the Fujisaki Model for F0 contour for the Portuguese Language. Phrase commands location in text is governed by a set of weighted rules. The amplitude (Ap) and timing (T0) of the phrase commands are predicted in separate neural networks. The features for both neural networks are discussed. Finally a comparison between target and predi...
متن کاملThis is a placeholder. Final title will be filled later
Solutions proposed in bibliography for multiple user allocation on the sub-bands of an OFDM system, adopting multiple antennas, require highly computational effort and consider delay insensitive applications. Our approach tends to overcome all these limitations relaxing some hypothesis in order to give a feasible solution. The proposed algorithm can be applied to a real multiple antenna OFDM sy...
متن کاملTODO: This is a placeholder. Final title will be filled later
Classification performance for emotional user states found in the few realistic, spontaneous databases available is as yet not very high. We present a database with emotional children’s speech in a human-robot scenario. Baseline classification performance for seven classes is 44.5%, for four classes 59.2%. We discuss possible strategies for tuning, e.g., using only prototypes (based on annotati...
متن کاملTODO: This is a placeholder. Final title will be filled later
We report work on mapping the acoustic speech signal, parametrized using Mel Frequency Cepstral Analysis, onto electromagnetic articulography trajectories from the MOCHA database. We employ the machine learning technique of Support Vector Regression, contrasting previous works that applied Neural Networks to the same task. Our results are comparable to those older attempts, even though, due to ...
متن کاملThis is a placeholder. Final title will be filled later
Recent auditory physiological evidence points to a modulation frequency dimension in the auditory cortex. This dimension exists jointly with the tonotopic acoustic frequency dimension. Thus, audition can be considered as a relatively slowly-varying two-dimensional representation, the “modulation spectrum,” where the first dimension is the well-known acoustic frequency and the second dimension i...
متن کاملThis is a placeholder. Final title will be filled later
Sine-wave speech (SWS) is a three-tone replica of speech, conventionally created by matching each constituent sinusoid in amplitude and frequency with the corresponding vocal tract resonance (formant). We propose an alternative technique where we take a high-quality multicomponent sinusoidal representation and decimate this model so that there are only three components per frame. In contrast to...
متن کامل